Skip to content

fix: parse truth tests after predicates - #2762

Open
JimmyLiuJia wants to merge 1 commit into
JSQLParser:masterfrom
JimmyLiuJia:fix/truth-tests-after-predicates
Open

JimmyLiuJia wants to merge 1 commit into
JSQLParser:masterfrom
JimmyLiuJia:fix/truth-tests-after-predicates

Conversation

@JimmyLiuJia

Copy link
Copy Markdown

Fixes #2759.

Statements such as a > 0 IS TRUE, a IN (1, 2) IS FALSE, and a IS NULL IS NOT TRUE currently fail at the final IS: the condition suffix dispatch treats truth tests as alternatives to the comparison or other predicate. Consume the optional IS [NOT] TRUE/FALSE/UNKNOWN after the complete predicate, reusing the existing expression types and preserving the precedence of leading NOT and outer AND/OR. A shared IS [NOT] prefix avoids separate Boolean and UNKNOWN lookahead probes. Bare and parenthesized truth tests continue to use the same AST types.

Add 25 regression/compatibility cases covering parse/deparse, AST structure, comparisons, IN, IS NULL, BETWEEN, LIKE, EXISTS, UNKNOWN, negation, logical precedence, select items, malformed suffixes, and the existing partial-expression parsing behavior. Thirteen regression cases fail on the unmodified base commit; all 25 cases pass with the fix.

Validation on Windows with JDK 17.0.20.1 and Gradle 9.8.0:

  • Full check passed, including grammar ambiguity, Checkstyle, PMD, SpotBugs, formatting, and coverage checks. JUnit XML reports 9,765 passed, 25 skipped, and no failures. Tests used a writable local temporary directory through java.io.tmpdir.
  • All three SQL statements from the report executed on SQLite 3.50.4 and returned the same rows as their explicitly parenthesized equivalents.
  • Official JSQLParserBenchmark.parseSQLStatements, version=latest, SIMPLE configuration, identical dependency classes and corpus, with a fixed 1 GiB heap and the GC profiler: 3 forks, 2 warmup iterations of 10 seconds, and 5 measurement iterations of 1 second per fork. Runs used baseline/fixed/fixed/baseline order.
Benchmark build Average time, ms/op (JMH 99.9% interval)
Before, run 1 76.876 +/- 6.180
After, run 1 91.153 +/- 39.118
After, run 2 74.405 +/- 3.883
Before, run 2 74.941 +/- 5.552

The first fixed run includes an unexplained slower fork (maximum measurement 208.510 ms/op); the repeated fixed run is comparable to the baseline runs. Keeping all measurements, the combined mean is 9.1% higher and the median is 1.9% lower. Allocation is approximately 7.356 million B/op in all four runs. These measurements cover the project's performance.sql corpus and cannot rule out smaller performance changes or changes in tail latency.

@manticore-projects

Copy link
Copy Markdown
Contributor

a IN (1, 2) IS FALSE can anyone explain to me please, what this is good for and why it should be supported?!

@JimmyLiuJia

Copy link
Copy Markdown
Author

For that particular WHERE example, a NOT IN (1, 2) is an equivalent filter. The motivation here is to accept existing SQL from #2759: SQLite executes the unparenthesized form, while JSqlParser rejects it.

IS FALSE returns a definite true/false value, whereas NOT preserves an unknown (NULL) result. For example, I checked this on SQLite 3.50.4:

SELECT NULL IN (1, 2) IS FALSE, NULL NOT IN (1, 2);
-- 0, NULL

So the difference matters when returning the boolean value, even though it does not change which rows the reported WHERE clause selects. SQLite documents this in Boolean expressions and the IN/NOT IN result matrix.

The parser change keeps the existing truth-test AST types and accepts the three forms in #2759.

@fudianchn

fudianchn commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor

For the original example, a IN (1, 2) IS FALSE explicitly tests whether the IN predicate is false. With t(a) = {1, 3, NULL}, it selects only 3 in a WHERE clause. It selects the same rows as NOT (a IN (1, 2)); the difference is the expression's value when a is NULL: IS FALSE returns FALSE, while NOT (...) remains UNKNOWN. Supporting it lets the parser handle this valid SQLite spelling directly.

For the broader truth-test family, IS NOT TRUE also has a distinct filtering use: retaining rows where a condition is false or unknown. When a is NULL, a > 0 IS NOT TRUE evaluates TRUE, while NOT a > 0 evaluates UNKNOWN, so the filters select different rows.

-- SQLite 3.53, t(a) = {1, 0, NULL}
SELECT count(*) FROM t WHERE a > 0 IS NOT TRUE;  -- 2
SELECT count(*) FROM t WHERE NOT a > 0;          -- 1

That behavior is covered by SQLite's boolean expressions and IS/IS NOT semantics.

There is also a related parser limitation worth tracking: IS [NOT] NULL and IS DISTINCT FROM after an unparenthesized predicate still fail. Parsed on master (40ed289b) and this PR head (678414fb):

--                                                master   PR head
SELECT * FROM t WHERE a IS TRUE;                  ok       ok
SELECT * FROM t WHERE (a > 0) IS TRUE;            ok       ok
SELECT * FROM t WHERE a > 0 IS TRUE;              fail     pass
SELECT * FROM t WHERE a > 0 IS NOT NULL;          fail     fail
SELECT * FROM t WHERE a = 1 IS NULL;              fail     fail
SELECT * FROM t WHERE a IS NOT NULL IS NULL;      fail     fail
SELECT * FROM t WHERE a > 0 IS DISTINCT FROM b;   fail     fail

a IS NOT NULL and (a > 0) IS NOT NULL parse on both builds. The corresponding unparenthesized IN/LIKE/BETWEEN cases also fail at the trailing IS.

The shared cause is the single condition-suffix slot in Condition(): once the comparison or predicate consumes it, there is no production for another trailing IS. This PR adds a separate truth-test suffix after the complete predicate; the null and distinctness tests remain in the original slot.

These are relevant dialect forms. PostgreSQL places IS tests below comparison operators, and MySQL's expression grammar allows boolean_primary IS [NOT] NULL, including after a comparison. These are syntax observations; execution still requires compatible operand types. For example, PostgreSQL's bare a IS TRUE requires a boolean a.

The PR improves the cases shown, with no regression in this comparison. I would prefer to cover the related null and distinctness families together. If that would broaden this PR too much, the remaining cases should be tracked explicitly as follow-up work.

@manticore-projects

Copy link
Copy Markdown
Contributor

Condition() is TRUE (or FALSE) makes zero sense in my opinion (no matter if SQLite allows it or not).
SELECT * FROM t WHERE a > 0 IS DISTINCT FROM b; is the only use-case that I see as somehow relevant (although I don't like it yet, we buy too much complexity without gaining much). So if anyone wants to write a complete PR, I would accept it. But please, everyone understand that I score the performance and simplicity of the parser high.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] JSQLParser 5.4 : PostgreSQL / SQLite : IS TRUE / IS FALSE after a comparison or predicate is a parse error

3 participants